Papers with crowdsourcing platform

8 papers
CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects (L18-1)

Copied to clipboard

Challenge: Various corpora of dialects have been collected using a well-equipped recording environment due to geographical and expense issues.
Approach: They construct a crowdsourced parallel speech corpus of Japanese dialects using crowdsourcing platforms.
Outcome: The proposed corpus includes parallel text and speech data of 21 Japanese dialects.
Crowdsourcing-based Annotation of the Accounting Registers of the Italian Comedy (L18-1)

Copied to clipboard

Challenge: CIRESFI project aims to reassess a theatrical heritage that has often been considered inferior to that of the two major, royally-privileged theaters.
Approach: They propose a double annotation system for new handwritten historical documents . crowdsourcing platform is set up to perform labeling and transcription of the documents based on budget data .
Outcome: The proposed system is based on a database of 25,250 pages of registers of the Italian Comedy of the 18th century.
Translation Crowdsourcing: Creating a Multilingual Corpus of Online Educational Content (L18-1)

Copied to clipboard

Challenge: a large corpus of online content has been developed via large-scale crowdsourcing.
Approach: They describe a multilingual corpus of online content that has been manually translated into 11 European and BRIC languages using the crowdsourcing platform.
Outcome: The proposed corpus is a product of the EU-funded TraMOOC project and is used to train, tune and test machine translation engines.
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Existing work on document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes.
Approach: They assemble and publish a multilingual Twitter corpus for the task of hate speech detection using inferred author demographic factors.
Outcome: The results show that the classifiers learn human biases and can be discriminatory towards certain demographic groups.
Strategies and Challenges for Crowdsourcing Regional Dialect Perception Data for Swiss German and Swiss French (L18-1)

Copied to clipboard

Challenge: a crowdsourcing project in the field of Swiss German dialects and Swiss French accents collects linguistic data.
Approach: a gamified crowdsourcing platform was set up to collect linguistic data on Swiss German and Swiss French accents.
Outcome: a gamified crowdsourcing platform collects linguistic data on Swiss German and Swiss French accents . the platform has provided 470,000 localizations, with 7,500 registered users and 30,000 anonymous visitors .
Nonsense!: Quality Control via Two-Step Reason Selection for Annotating Local Acceptability and Related Attributes in News Editorials (D19-1)

Copied to clipboard

Challenge: Annotation quality control is critical for building reliable corpora through linguistic annotation.
Approach: They propose a method to control annotation quality using two-step reason selection using a crowdsourcing platform.
Outcome: The proposed method retains the annotations with satisfactory quality out of the entire annotations mixed with those of low quality.
Rigor Mortis: Annotating MWEs with a Gamified Platform (2020.lrec-1)

Copied to clipboard

Challenge: gamification of the platform should be improved, in order to attract and retain more players.
Approach: They propose to use a gamified crowdsourcing platform to evaluate the intuition of speakers and then train them to annotate multi-word expressions in French corpora.
Outcome: The proposed platform evaluates the speakers' intuition and trains them to annotate multi-word expressions in French corpora.
A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection (2020.lrec-1)

Copied to clipboard

Challenge: Social media platforms allow users to engage in conversation with limited accountability, causing hate crimes and mental harm to targeted individuals.
Approach: They propose to make public a new dialectal Arabic news comment dataset . they analyze distinctive lexical content along with the use of emojis in offensive comments .
Outcome: The proposed dataset analyzes offensive language and distinctive lexical content along with the use of emojis on Twitter, Facebook, and YouTube.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations